Skip to content

Add repeating character preprocessor - #2444

Open
ShreyasB1 wants to merge 4 commits into
gunthercox:masterfrom
ShreyasB1:add-repeating-character-preprocessor
Open

ShreyasB1 wants to merge 4 commits into
gunthercox:masterfrom
ShreyasB1:add-repeating-character-preprocessor

Conversation

@ShreyasB1

Copy link
Copy Markdown

Add a preprocessor that reduces runs of three or more repeated letters
down to two (e.g. "sooooo" -> "soo"). Elongated words are common in
conversational input, and normalizing them maps spelling variations to a
consistent form so the chat bot can match input against trained
statements more reliably.

Repeated digits and punctuation are left unchanged, and naturally
occurring double letters (such as the "oo" in "cool") are preserved.

Includes unit tests and documentation for the new preprocessor.

@jocelyn1981

jocelyn1981 commented Jun 7, 2026 via email

Copy link
Copy Markdown

@ShreyasB1

Copy link
Copy Markdown
Author

Hi! I was wondering the timeline for this feature being approved.

Comment thread chatterbot/preprocessors.py Outdated
preserved, and repeated digits or punctuation are left unchanged.
"""
statement.text = _REPEATING_CHARACTER_PATTERN.sub(
lambda match: match.group(1) * 2, statement.text

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Taking a look at some related sample data, it appears that this regex would group a statement such as "I am sooooo happy" with "I am soo happy" (x2 characters), which still doesn't quite reach the intended token of "so".

I'm not certain this works as intended in some of the cases the pull request was expecting. Perhaps there is another approach that might work better? (If not, a project-specific preprocessor is always an alternative option to including one in the main codebase.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch — you're right, and I reworked it.

I went looking for a way to tell "sooooo" → "so" apart from "gooood" → "good", and the honest answer is that it needs a dictionary. I checked whether spaCy's vocab could serve as one, and it can't: with en_core_web_sm, "soo" and "col" both report as present in nlp.vocab.strings, and is_oov is True for every token because the small model ships no vectors. Adding a real word list felt out of scope for a preprocessor.

So instead of reducing the run to two characters, it's now reduced to one:

  • "I am sooooo happy" → "I am so happy"
  • "Yesss that was greaaaat" → "Yes that was great"

This is safe because the pattern only matches runs of three or more, and no correctly spelled English word has one. Checking against /usr/share/dict/words (235,976 entries), only 7 match, all archaic -ssship constructions (bossship, goddessship, wallless). So the preprocessor never alters text that was already spelled correctly. Naturally doubled letters are runs of two, so "cool" is still untouched — the two-character version was protecting a case that was never at risk.

The tradeoff that remains: a word that genuinely contains a doubled letter is reduced past its correct spelling when elongated, so "gooood" → "god". That's documented in the docstring and pinned by a test, with a note that a project needing the distinction can register its own preprocessor backed by a word list.

Also fixed the flake8 errors (E305/E302/W293) that were in my earlier commits.

Happy to close this and keep it project-specific instead if you'd still rather not carry the ambiguity in core.

@syedkosaainhaider-maker

Copy link
Copy Markdown

ok

Reducing a run of three or more repeated letters down to two does not
reach the word being elongated: "I am sooooo happy" became "I am soo
happy" rather than "I am so happy", so elongated input still did not
group with the trained statement it was meant to match.

Reduce the run to a single character instead. Because no correctly
spelled English word contains the same letter three or more times in a
row, only text that was already non-standard is altered, and naturally
doubled letters are still untouched because they are runs of two
("cool" is unchanged).

A word that genuinely contains a doubled letter is now reduced past its
correct spelling when elongated ("gooood" -> "god"). Telling that case
apart from "sooooo" -> "so" requires a dictionary lookup, which is out
of scope here; this is documented in the docstring and covered by a
test.

Also fix the flake8 errors in the previous commits (E305, E302, W293).
@jocelyn1981

jocelyn1981 commented Sep 21, 2026 via email

Copy link
Copy Markdown

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants